Agentic AI in MES: Who Decides What the Software Is Allowed to Do

Manufacturing control room screens showing production scheduling and quality dashboards

For the past couple of years, “AI in MES” mostly meant a chat window bolted onto a dashboard. Ask it why a line is down, ask it to summarize yesterday’s scrap, ask it which work centers are trending toward a quality excursion. Useful, low-risk, and — critically — it left a human in the loop before anything actually happened on the floor. That era isn’t over, but it’s being layered with something more consequential: agents that don’t just answer the question, they act on the answer. Reschedule the order. Release the next lot. Put a hold on a work center. Kick off a replenishment transaction against a supplier’s punchout catalog.

That shift — from advisory copilot to executing agent — is the real story heading into 2026, and it’s happening faster than most plants’ change-management processes can absorb. The mistake I’d flag early: treating this as a rollout decision when it’s actually a governance decision. The question isn’t “should we turn on the agentic scheduling feature.” It’s “what, specifically, is this thing allowed to touch without a human confirming it first” — and if you don’t have an answer before the vendor flips the switch, you’ll get one by accident.

What’s actually new here

Traditional MES automation has always executed things without asking — that’s what a dispatch rule or an interlock is. The difference with agentic AI is that the logic isn’t a fixed rule anymore; it’s a model reasoning over context, and the action space is broader and less predictable by design. An agent that manages scheduling might reprioritize dozens of orders based on a changed due date, a machine going down, or a material shortage signal from an ERP integration — chaining several decisions together without a person reviewing each one. That’s meaningfully different from a scheduling engine executing a rule someone wrote and validated. Nobody fully validated the agent’s reasoning path in advance, because that’s the point of the thing.

This is also arriving at the same time MES platforms are getting deeper native connections into scheduling, quality, and inventory modules — the same surfaces that used to require a human to click “release” or “approve hold.” Vendors are marketing this as autonomy. It is autonomy. The industry has decades of experience with autonomous control at the equipment layer, governed by things like IEC 62443 for security and well-worn interlock and permissive logic for safety. We do not yet have an equivalent, widely adopted discipline for autonomous decision-making at the MES layer. That gap is where the risk lives.

A tiering model, before you turn anything on

Rather than treating “agentic” as one feature you enable or disable, it helps to think of every MES function as sitting on a spectrum of autonomy, and to explicitly decide — function by function — which tier it’s allowed to occupy.

Tier 0: Read-only

The agent observes and reports. It can flag an anomaly, summarize a trend, or answer a natural-language query against production data. It cannot write anything back to the MES. This is where most 2024–2025 copilots have lived, and it’s still the right tier for anything where a wrong answer costs you a few minutes of skepticism rather than a bad batch.

Tier 1: Propose-with-approval

The agent generates a specific, reviewable recommendation — a reordered schedule, a suggested hold, a proposed substitute material — and a named human has to approve it before it executes. This is the tier where agentic AI earns real trust: the agent does the cognitive heavy lifting, a person with authority and context does the sign-off. Scheduling changes that affect customer due dates, and any quality disposition that could release nonconforming material, belong here at minimum until you have a long track record of the agent being right.

Tier 2: Bounded autonomy

The agent acts without a human in the loop, but only inside explicit, narrow limits you’ve set — a quantity cap, a value threshold, a defined set of SKUs, a specific work center. Routine replenishment triggers for non-critical MRO or indirect materials are a reasonable candidate: let the agent auto-generate a purchase requisition when inventory crosses a reorder point, within a spending limit, for a pre-approved supplier list. Minor schedule sequencing within a single shift, among orders that are already released and don’t touch due-date commitments, is another plausible candidate. The discipline here is that “bounded” has to mean something enforced by the system, not just written in a policy document — hard limits, not guidelines the agent is asked nicely to respect.

Tier 3: Full autonomy

The agent acts, and nobody reviews it before or immediately after. In our assessment, almost nothing on a production floor belongs here yet — not quality holds, not schedule releases that touch customer commitments, not anything with safety or regulatory exposure. Full autonomy is where vendor demos live and where most plants should not follow, at least not in 2026.

Mapping functions, not features

The practical exercise is to go through your MES module by module and assign a tier, in writing, with a named owner who can change it. Quality holds and dispositions should almost never exceed Tier 1 — the cost of a false release is too asymmetric to the cost of a delayed one. Scheduling is more nuanced: sequencing within already-released work can tolerate more autonomy than anything that reprioritizes across customer orders or touches capacity commitments. Material replenishment is probably your best early candidate for Tier 2, because the failure mode — over-ordering a commodity item — is cheap relative to a missed quality catch. Work order release, especially for anything that consumes serialized or lot-controlled material, deserves the same caution as quality.

Two things make this tiering enforceable rather than aspirational. First, every agent action, at every tier above 0, needs an audit trail that’s as rigorous as an electronic batch record — what triggered it, what data it reasoned over, who approved it or that no one did. Second, you need a fast, well-understood way to drop any function back a tier or to zero, immediately, without a change request wending through IT. Agentic features fail in ways rule-based automation doesn’t: not by breaking loudly, but by making a plausible-sounding call on ambiguous or novel input that a rigid rule would have simply refused to touch.

What to actually do this year

Don’t wait for the vendor’s roadmap to force the conversation. Build the tiering document now, before your MES contract renewal or upgrade cycle puts an “enable agentic scheduling” toggle in front of you. Get quality, production control, and plant IT in the same room to agree on tiers per function — this is not a decision controls engineering or IT should make alone, and it’s not one procurement should make by accepting a vendor’s default configuration. Ask any vendor pitching agentic capability exactly what tier each feature defaults to, whether that default is configurable per function, and what the audit trail looks like. If they can’t answer precisely, that’s your answer for now.

The technology is real and the direction is right — plenty of scheduling and replenishment decisions are genuinely better handled by a model reasoning over live data than by a static rule table. But autonomy earned gradually, with a paper trail and a kill switch, beats autonomy granted by default because a demo looked impressive.


This article was written with the assistance of artificial intelligence. While we aim for accuracy, the information may be incomplete, out of date, or incorrect, and should be independently verified before you rely on it for any decision. It is provided for general information only and does not constitute professional advice.

Related posts